video story
Facilitating Video Story Interaction with Multi-Agent Collaborative System
Zhang, Yiwen, Hao, Jianing, Wang, Zhan, Sheng, Hongling, Zeng, Wei
Video story interaction enables viewers to engage with and explore narrative content for personalized experiences. However, existing methods are limited to user selection, specially designed narratives, and lack customization. To address this, we propose an interactive system based on user intent. Our system uses a Vision Language Model (VLM) to enable machines to understand video stories, combining Retrieval-Augmented Generation (RAG) and a Multi-Agent System (MAS) to create evolving characters and scene experiences. It includes three stages: 1) Video story processing, utilizing VLM and prior knowledge to simulate human understanding of stories across three modalities. 2) Multi-space chat, creating growth-oriented characters through MAS interactions based on user queries and story stages. 3) Scene customization, expanding and visualizing various story scenes mentioned in dialogue. Applied to the Harry Potter series, our study shows the system effectively portrays emergent character social behavior and growth, enhancing the interactive experience in the video story world.
StoryLine: Exploring the intersection of visual storytelling and machine learning โ MIT Media Lab
StoryLine, a collaboration between McKinsey & Company's Consumer Tech and Media team and the Lab for Social Machines (LSM), explores the intersection of visual storytelling and machine learning through work aimed at helping storytellers understand and improve the impact of their stories on their audiences. StoryLine was inspired by LSM's groundbreaking Electome project, in which researchers used advanced machine learning to build network maps of engaged election audiences and then track the diffusion of relevant conversation and content through these networks. For Electome, the result was a powerful new way of understanding how audiences form around specific political, social and cultural ideas. For StoryLine, we wondered whether this approach could translate meaningfully into measuring impact in the storytelling domain, which outside of marketing optimization has been underserved in any practical way by recent advances in machine learning and artificial intelligence. Understanding audience impact has arguably never been more important to those involved in the creation, production, distribution and marketing of visual stories--from the largest studios, networks and platforms to the independent creators who publish and promote on their own.
Converting text news into video stories using Deep Learning
What these companies have in common is that they both use cutting-edge deep learning technologies such as generative adversarial networks, facial point detection, body pose estimation, and cloud technologies to synthesize human faces and videos that cannot be distinguished from reality. Today, I want to show and deconstruct a new product in the synthetic media category that we have been developing for the last few months called NIUS.TV. Catching up on news on mobile is still painful. Reading while commuting, exercising, or waiting is difficult. NIUS.TV is a next-generation mobile-first news aggregator that converts text news on topics you love into video stories narrated by an AI anchor.
DramaQA: Character-Centered Video Story Understanding with Hierarchical QA
Choi, Seongho, On, Kyoung-Woon, Heo, Yu-Jung, Seo, Ahjeong, Jang, Youwon, Lee, Seungchan, Lee, Minsu, Zhang, Byoung-Tak
Despite recent progress on computer vision and natural language processing, developing video understanding intelligence is still hard to achieve due to the intrinsic difficulty of story in video. Moreover, there is not a theoretical metric for evaluating the degree of video understanding. In this paper, we propose a novel video question answering (Video QA) task, DramaQA, for a comprehensive understanding of the video story. The DramaQA focused on two perspectives: 1) hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence. 2) character-centered video annotations to model local coherence of the story. Our dataset is built upon the TV drama "Another Miss Oh" and it contains 16,191 QA pairs from 23,928 various length video clips, with each QA pair belonging to one of four difficulty levels. We provide 217,308 annotated images with rich character-centered annotations, including visual bounding boxes, behaviors, and emotions of main characters, and coreference resolved scripts. Additionally, we provide analyses of the dataset as well as Dual Matching Multistream model which effectively learns character-centered representations of video to answer questions about the video. We are planning to release our dataset and model publicly for research purposes and expect that our work will provide a new perspective on video story understanding research.
AI in storytelling: Machines as cocreators
Sunspring debuted at the SCI-FI LONDON film festival in 2016. Set in a dystopian world with mass unemployment, the movie attracted many fans, with one viewer describing it as amusing but strange. But the most notable aspect of the film involves its creation: an artificial-intelligence (AI) bot wrote Sunspring's screenplay. "Maybe machines will replace human storytellers, just like self-driving cars could take over the roads." A closer look at Sunspring might raise some doubts, however.
AI in storytelling: Machines as cocreators
Computers don't cry during sad stories, but they can tell when we will. Sunspring debuted at the SCI-FI LONDON film festival in 2016. Set in a dystopian world with mass unemployment, the movie attracted many fans, with one viewer describing it as amusing but strange. But the most notable aspect of the film involves its creation: an artificial-intelligence (AI) bot wrote Sunspring's screenplay. "Maybe machines will replace human storytellers, just like self-driving cars could take over the roads."